Papers with multi-document summarization task
Multi-News+: Cost-efficient Dataset Cleansing via LLM-based Data Annotation (2024.emnlp-main)
Copied to clipboard
| Challenge: | Various attempts to correct noisy data in the construction process have been made, but human annotation is expensive and time-consuming. |
| Approach: | They propose to use large language models for data annotation to imitate human annotation and classify unrelated documents from a multi-document summarization task. |
| Outcome: | The proposed method imitates human annotation and classifies unrelated documents from the Multi-News dataset. |
Unsupervised Aspect-Based Multi-Document Abstractive Summarization (D19-54)
Copied to clipboard
| Challenge: | Existing methods for opinion summarization are expensive and do not deal with contradictory statements. |
| Approach: | They propose an unsupervised abstractive summarization neural system that generates short summaries of reviews in a vector space. |
| Outcome: | The proposed system can generate short summaries of user-generated reviews in a short paragraph, while nobody reads all reviews. |
What’s in a Summary? Laying the Groundwork for Advances in Hospital-Course Summarization (2021.naacl-main)
Copied to clipboard
| Challenge: | Existing methods to summarize clinical narratives are lacking. |
| Approach: | They propose to generate a paragraph that tells the story of a patient's hospitalization . they analyze a dataset of 109,000 hospitalizations and their corresponding summary proxy . |
| Outcome: | The proposed model is based on a dataset of 109,000 hospitalizations and their corresponding summary proxy. |
Multi-XScience: A Large-scale Dataset for Extreme Multi-document Summarization of Scientific Articles (2020.emnlp-main)
Copied to clipboard
| Challenge: | Multi-XScience is a dataset construction protocol that favours abstractive modeling approaches. |
| Approach: | They propose a large-scale multi-document summarization dataset that is based on articles and lexical databases and WordNet synonymy information to generate related-work sections of a paper. |
| Outcome: | The proposed method is based on lexical databases and WordNet synonymy information to write related work sections of a paper based upon their abstract and the articles they reference. |
Benchmarking LLMs on Semantic Overlap Summarization (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are the most capable text generation models in a variety of tasks and fields. |
| Approach: | They benchmark Large Language Models (LLMs) on SOS and introduce PrivacyPolicyPairs (3P) a dataset of 135 high-quality privacy policy documents is used to evaluate the model. |
| Outcome: | The proposed dataset complements existing resources and broadens domain coverage. |